Back

Journal of Genetics and Genomics

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Journal of Genetics and Genomics's content profile, based on 38 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Using CRISPR/Cas9 to investigate the role of candidate human disease gene orthologs in Ciona

Hernandez, S. A.; Johnson, C. J.; Stolfi, A.

2026-08-11 developmental biology 10.64898/2026.08.10.743552 medRxiv
Top 0.1%
5.1%
Show abstract

The tunicate Ciona robusta offers a tractable non-vertebrate chordate model for probing gene function via tissue-specific, CRISPR/Cas9-mediated mutagenesis in F0. Building on Arcadia Sciences Zoogle platform, which identifies and ranks orthologs of human genes from various non-traditional model organisms, we carried out a pilot project to probe the developmental roles of three notochord- and endoderm-expressed candidate orthologs of human disease genes (Fcho, Pgm3, and Nckap1) alongside a fourth gene (Plastin) implicated in papilla cell elongation. This preprint compiles and updates a series of research project milestones previously posted episodically on Zenodo. Here we summarize the full results and our conclusion about this pilot project. Using CRISPR/Cas9, we found that tissue-specific knockout of Pgm3 and, to a lesser extent, Fcho caused significant defects in larval tail elongation. Separately, CRISPR knockout of Plastin, an actin-bundling gene expressed throughout the sensory-adhesive papillae of the larva, caused a subtle reduction in papilla cell elongation when combined as a duoble knockout with another actin-bundling protein-encoding gene, Villin. These results identify Pgm3 as the most promising candidate for further development as a Ciona-based model of human disease and demonstrate the utility of tissue-specific CRISPR screening for prioritizing candidate disease gene orthologs identified through comparative genomics platforms like Zoogle.

2
u4atac regulates cilium biogenesis through splicing of the minor intron of tmem107l and rfx7b in zebrafish developing brain

Jovani, C.; Rabec, A.; Gaubert, M.; Khatri, D.; Garnier, E.; Cologne, A.; Meiller, A.; Guguin, J.; Besson, A.; Mazoyer, S.; DELOUS, M.

2026-08-24 genetics 10.64898/2026.08.20.745718 medRxiv
Top 0.1%
4.3%
Show abstract

Bi-allelic variants of RNU4ATAC, transcribed into the minor spliceosome component U4atac snRNA, are associated to variable severity of microcephaly, growth retardation, skeletal dysplasia and immunodeficiency as main features. Previous studies highlighted the dramatic effect of U4atac deficiency on splicing of U12-type introns, which represent less than 1% of all introns in the human genome. More recently, our team evidenced a link between U4atac and the primary cilium/centrosome complex through the identification of patients carrying RNU4ATAC bi-allelic variants and exhibiting an atypical Joubert syndrome, a well-known ciliopathy. Yet, the underlying mechanisms remain elusive. Here, we further explored the link of RNU4ATAC to primary cilium and aimed at identifying ciliary U12-type intron containing genes that contribute to the brain abnormalities seen in patients. For that, we performed a transcriptomic analysis of heads of our morpholino oligonucleotide (MO)-mediated u4atac zebrafish model. Through the combined analysis of the generated dataset with those obtained from RNU4ATAC patient cells, we identified two candidate genes: TMEM107, coding for a structural protein of the cilium transition zone, and RFX7, encoding a transcription factor involved in primary cilium formation. By conducting complementary genetic approaches in zebrafish model, we showed that both gene orthologues, tmem107l and rfx7b, functionally interact with u4atac and are required for correct brain development. Altogether, our findings establish TMEM107 and RFX7 as key components of the molecular pathway linking U4atac dysfunction to ciliary defects and impaired brain development, providing new physiopathological insights and therapeutic perspectives for RNU4ATAC-related disorders.

3
Port of Protein-Protein Interactomes: An experiment-based protein-protein interactome database for rice

Liu, X.; Lu, J.; Jia, L.; Xia, D.; Huang, J.; Cheng, Y.; Li, M.; Chen, Y.; Liu, X.; Li, G.; Liu, W.; Li, J.; Ying, J.; Wang, Y.; Li, Z.; Tong, X.; Hou, Y.; Zhiguo, E.; Zhang, J.; Zhang, J.

2026-08-20 systems biology 10.64898/2026.08.16.744343 medRxiv
Top 0.1%
3.2%
Show abstract

Protein-protein interactions (PPIs) play a crucial role in enabling proteins to carry out their functions within various biological processes (Hui et al., 2003). Since the introduction of the yeast two-hybrid (Y2H) method for PPI detection in 1989 (Fields and Song, 1989), the identification of PPIs has become a significant focus in modern biological research. PPI goes beyond examining individual proteins, allowing researchers to establish a comprehensive network that regulates biological processes. Rice, as a key model organism in plant biological studies, has been at the forefront of PPI research. In 2008, prominent rice scientists in China called for concerted efforts to define a comprehensive protein-protein interaction network experimentally, which aimed to facilitate the prediction of the functional mechanisms operating throughout a plants lifecycle (Zhang et al., 2008). With efforts for 2 decades, the experimentally identified rice PPIs have reached over ten thousand. Several public databases have been established to systematically collate and store PPIs, including STRING (Szklarczyk et al., 2019), BioGRID (Oughtred et al., 2020), IntAct (del Toro et al., 2022), PRIN (Gu et al., 2011), RicePPINet (Liu et al., 2017) and RiceNet v2 (Lee et al., 2015). However, most PPI datasets in rice stem from computational predictions, while experiment-based rice PPI datasets are fragmented due to the lack of systematic profiling at the rice PPIome level, which largely hinders information sharing in the rice research community. To bridge this gap, we constructed the Port of Protein-Protein Interactomes (POPPIN; https://riceome.hzau.edu.cn/poppin/), an integrated database dedicated to sharing experimentally verified PPIs and functional clues in rice. Empowered by high-throughput PPIome profiling technologies and text mining assisted by a large language model (Huang et al., 2025; Liu et al., 2025), POPPIN currently has deposited over 150,451 pieces of rice PPI-related information. Additionally, POPPIN provides detailed protein information, including GO annotations, subcellular localizations, domains, trait ontology (TO) information, and hyperlinks to external biological databases. Through offering a user-friendly web interface for search and dynamic network visualization, POPPIN serves as the first large-scale, experiment-based database for searchable PPIs in rice, and has the potential to be extended to other species under this structural framework.

4
A Practical Framework for Constructing Population-Specific and Alternate-Contig-Aware Genome References: A case study of Vietnam

Vo, N. S.; Tran, T. T. H.; Duong, V. C.; Nguyen, N. N.; Pham, T. M.; Vu, Q. T.; Tran, M. H.; Hoang, T. H.; Nguyen, Q.; Nguyen, D. T.

2026-08-27 genomics 10.64898/2026.08.24.746817 medRxiv
Top 0.2%
2.0%
Show abstract

Current studies in human genomics typically rely on the standard genome reference GRCh38 which is known to be biased toward populations of European ancestry and therefore has limitations when applied to other populations. Although various graph-based pangenome references were constructed for several populations to deal with this bias, their usage in practice is currently still limited compared to linear genome references. Here we present a framework for constructing a population-specific genome reference using GRCh38 as backbone with alternate-contig awareness to enhance genomic data analysis in the target population. We demonstrated the advantages of our framework using both public and in-house Vietnamese whole-genome sequencing (WGS) datasets. Genomic variants derived from high-coverage WGS data of the 1000 Vietnamese Genomes Project (VN1K) were imported into our framework to build a Vietnamese-specific Genome Reference (VGR). VGR was then compared to GRCh38 in read alignment and variant calling using high-coverage WGS data of 99 Vietnamese individuals (KHV) from the 1000 Genomes Project (1kGP). Using Omni array genotyping data from 99 KHV samples as an independent benchmark, we found that VGR improved variant-calling precision and reduced false-positive calls compared to GRCh38. Our framework could be easily used for other populations as long as they have a variant database similar to VN1K. Our code is publicly available at github.com/VinGenome/VGR

5
Exploring vulnerable proteins in the progression of head and neck squamous cell carcinoma

Agrawal, A.; Kumar, S.; Vindal, V.

2026-08-13 bioinformatics 10.64898/2026.08.07.743269 medRxiv
Top 0.3%
1.7%
Show abstract

A protein whose removal or deletion causes significant disruption or collapse of a protein-protein interaction (PPI) network is referred to as a vulnerable protein. Such proteins may serve as valuable therapeutic or diagnostic targets in disease-associated networks. In this study, two PPI networks were constructed, one for HPV-positive and the other for HPV-negative head and neck squamous cell carcinoma (HNSCC), and the vulnerable proteins of these networks were identified by the node deletion approach. After analyzing the networks, 27 unique vulnerable proteins in HPV-positive and 72 unique vulnerable proteins in HPV-negative HNSCC were identified. Among them, one HPV-positive and seven HPV-negative HNSCC vulnerable proteins were further chosen by integrating multi-omics data. To exploit the vulnerabilities of these proteins, candidate synthetic lethal (SL) partners were predicted whose inhibition may selectively impair tumor survival. Subsequently, drug-gene interaction analysis was performed to identify inhibitors targeting the SL partners of these vulnerable proteins. Notably, in HPV-positive HNSCC, TOP2A, CHEK1, and CHEK2 genes were identified as SL partners of TTN, and their inhibitors were already clinically approved. While in HPV-negative HNSCC, ADA and MMP19 were identified as an SL partner of LMO7; TMEM45B, CDH3, and ELF3 genes were identified as an SL partner of CGN; and ZNF433 was identified as an SL partner of FLNC. However, MMP19, ZNF433, and TMEM45B inhibitors were not reported. Thus, these vulnerable proteins, including their SL partners, provide novel avenues to explore and develop more efficient and precise therapeutic and diagnostic strategies.

6
Multimodal spatial-omics reveal the heterogeneity and intercellular network characteristics of papillary craniopharyngiomas.

Jiang, Y.; Luo, H.; Zheng, H.; Li, C.; Zan, X.; Xu, J.; Chen, Y.

2026-08-24 cancer biology 10.64898/2026.08.20.746031 medRxiv
Top 0.3%
1.7%
Show abstract

Despite significant advancements in microsurgical techniques in recent years, the treatment and prognosis of craniopharyngiomas remain unsatisfactory. As a central nervous system tumor located adjacent to important brain structures such as the hypothalamus-pituitary axis and accompanied by a highly inflammatory microenvironment, the tumor heterogeneity and tumor microenvironment characteristics of papillary craniopharyngiomas (PCPs) remain unclear. In this study, we integrated multimodal single-cell and spatial profiling from PCP tissue and peripheral blood mononuclear cells (PBMCs) to elucidate the tumor heterogeneity and microenvironment characteristics of PCP. Our single-cell and spatial analyses defined four specific tumor cell states in PCP, representing specific transcriptional regulatory programs and spatial heterogeneity characteristics during tumor progression. By constructing a spatial niche composed of tumor, immune, and stromal cells, we analyzed the cellular and spatial ecosystem of PCP at multiple levels to further assess the communication relationships between different tumor cell states and microenvironment cells. This study established a multidimensional molecular atlas of PCP from the perspectives of cell state, spatial structure, and microenvironment interactions, providing a foundation for understanding its biological behavior and exploring new intervention strategies.

7
Ori-Finder-Arch: An Updated Web Server for the Annotation and Visualization of Archaeal Replication Origins

You, Z.; Zhang, Z.; Luo, H.; Gao, F.

2026-08-19 bioinformatics 10.64898/2026.08.15.744077 medRxiv
Top 0.5%
1.0%
Show abstract

Archaea are promising chassis organisms in biotechnology, and the accurate annotation of their chromosomal replication origins (oriCs) is the key to unlocking their full potential. However, the existing Ori-Finder 2 web server suffers from low accuracy, slow speed, and limited scalability. In this study, we present Ori-Finder-Arch, an updated web server for high-performance oriC prediction in archaea. This pipeline integrates HMMER-based replication initiation protein (RIP) annotation, refined consensus motif recognition, and GC profile-based DNA unwinding element (DUE) detection. On a benchmark set of experimentally validated oriCs, Ori-Finder-Arch achieved a recall of 95.6% and a precision of 86.0%, substantially outperforming Ori-Finder 2 (62.2% and 63.6%, respectively), while running 4.75 times faster and supporting diverse assembly levels. When applied to the available archaeal assemblies, it successfully annotated 17,472 oriCs. Meanwhile, the web server provides interactive visualizations at different levels. In conclusion, Ori-Finder-Arch offers an efficient, accurate, and user-friendly platform for advanced studies of archaeal DNA replication initiation and synthetic biology applications, and is freely available at https://tubic.org/Ori-Finder-Arch/ and https://tubic.tju.edu.cn/Ori-Finder-Arch/.

8
Long-read single cell transcriptomics uncovers isoform preferences in developing human retina

Kaplan, L.; Pang, J.; Reh, T. A.

2026-08-21 developmental biology 10.64898/2026.08.17.745278 medRxiv
Top 0.6%
1.0%
Show abstract

Retinal development has been extensively studied and key transcriptional regulators that drive fate decisions have been identified for major cell classes. These findings were confirmed and deepened in recent years with the advance of single cell RNA sequencing (Scrase). However, many processes that guide progenitor to postmitotic cell differentiation remain elusive, especially since some genes seem to yield different cell populations without apparent correlation with expression level or timing. Here, differential transcript isoform usage might play a role in diversifying the function of developmental genes. In short-read based scRNAseq, isoforms can only be identified if a read maps to a unique sequence or exon junction. However, due to the sparsity and very short reads, these events are extremely rare. We combined a commercial scRNAseq kit, that produces barcoded, full-length cDNA with Oxford Nanopore Technologies based long-read sequencing to generate the first single cell long-read sequencing dataset of fetal human retina. It can help elucidate the role of alternative splicing in retinal development and guide the design of transcript-specific gene therapies for retinal regeneration.

9
Generation of Human Taste Bud Organoids as a Human-Mimetic Platform for Modeling Taste Perception

Chae, J.; Kwon, S. S.; Kim, J.; Moon, H.; Do, V. Q.; Zehentner, S.; Cho, H.-J.; Bhin, J.; Moon, S. J.; Kim, C. H.

2026-08-10 developmental biology 10.64898/2026.08.09.743674 medRxiv
Top 0.7%
0.9%
Show abstract

We established a human taste bud organoid system derived from circumvallate papillae. This model has been highly anticipated in the field of taste research, where feasible approaches for validating taste biology discovered in rodent models have been limited. Through a stepwise exploratory strategy, we systematically identified and optimized the niche factors required to maintain taste bud organoids and promote their differentiation. This human taste bud organoid system comprises Type I-IV taste receptor cells (TRCs) as well as stem/progenitor cells, and its sensory receptor cells exhibit calcium responses to taste stimuli. Using this system, we identified robust Wnt signaling as a requirement for optimal TRC fate progression, uncovered a human-specific transcriptional program in LGR5 cells, and identified previously unrecognized molecular markers for Type I TRCs. By recapitulating native human taste bud cell diversity and function, this organoid provides a tractable platform for studying human taste biology and dysfunction.

10
Integrative optical genome mapping and long-read sequencing resolve constitutional complex rearrangements at nucleotide resolution

Burssed, B.; van der Sanden, B.; Hops, W.; Neveling, K.; Kamping, E.; van Beek, R.; den Ouden, A.; Derks, R.; Timmermans, R.; Perrone, E.; Ramos, M. A.; Bellucco, F. T.; Hoischen, A.; Melaragno, M. I.

2026-08-28 genomics 10.64898/2026.08.27.747510 medRxiv
Top 0.7%
0.8%
Show abstract

Complex rearrangements are one of the rarest types of structural variants (SVs) and can be divided into two categories: complex chromosomal rearrangements (CCRs) and complex genomic rearrangements (CGRs). CCRs include structural rearrangements that present at least three breakpoints and show exchange of genetic material between more than two chromosomes and CGRs are rearrangements that present more than one junction and/or more than one SV in cis. They are usually formed by one of the chromoanagenesis mechanisms, where a massive disruptive cellular event leads to multiple structural rearrangements. Classical cytogenomic techniques have been commonly applied for their characterization, but methodologies that involve longer DNA molecules, namely optical genome mapping (OGM) and long-read genome sequencing (lrGS), present a considerably higher SV detection resolution, revealing more details about the rearrangements, including precise breakpoint location. Here, we describe six patients with complex rearrangements investigated through a combination of different techniques: karyotyping, chromosomal microarray, and OGM were performed to characterize the rearrangements. Subsequently, lrGS was used to further resolve the alterations, refine their breakpoints' location, and sequence their junction points. Three patients presented CCRs involving three, four, and six chromosomes, while three exhibited CGRs involving one different chromosome each, providing a variety of complex SVs to show the importance of each technique and their combination in rearrangement resolution. In total, the complex rearrangements presented 127 breakpoints, 66 junction points and involved 14 of the 24 chromosomes. Higher-resolution techniques revealed additional complexity in all cases. Despite the advances provided by OGM and lrGS, conventional karyotyping remained indispensable for complete rearrangement resolution. In two patients, the findings supported a novel mechanism combining features of the different chromoanagenesis processes. Furthermore, evidence of inherited alterations was identified, and the comprehensive characterization of the rearrangements enabled more accurate genotype-phenotype correlations. Our findings indicate that an integrated approach combining karyotyping, OGM, and lrGS can completely resolve SVs, including complex rearrangements.

11
mV2G: a multiomic atlas for tissue-specific variant-to-gene prioritization

Zhang, D. Y.; Zhou, H.; Sheth, M. U.; Gschwind, A. R.; Engreitz, J. M.; Lin, X.; Liu, H.

2026-08-20 genomics 10.64898/2026.08.11.743997 medRxiv
Top 0.7%
0.8%
Show abstract

The multiomic Variant-to-Gene (mV2G, https://mv2g.hbliulab.org) is a comprehensive atlas that integrates diverse functional genomic evidence to prioritize tissue-specific variant-to-gene (V2G) associations. While genome-wide association studies (GWAS) have identified millions of associations between genetic variants and diseases, translating these findings into biological mechanisms remains challenging because >90% of variants reside in noncoding regions. Existing V2G resources provide complementary regulatory evidence but are fragmented and often lack tissue-specific interpretation. To address this challenge, we constructed the mV2G atlas by integrating 24 types of functional genomic evidence across 50 human tissues, including molecular quantitative trait loci, enhancer-gene predictions, three-dimensional chromatin interactions, and experimental validation. The atlas contains 188,634,118 evidence-supported V2G pairs involving 13,618,039 variants and 69,521 genes. We further developed a unified tissue-specific V2G prioritization framework and prioritized 1,530,420 high-confidence functional V2G pairs involving 1,131,316 unique variants, with 87% exhibiting tissue-specificity. The mV2G atlas provides searchable variant- and gene-centered interfaces, an interactive browser for visualizing variants, target genes, cis-regulatory elements, and chromatin states, as well as downloadable datasets. By integrating complementary regulatory evidence into a unified framework, mV2G provides an accessible resource for interpreting the functional and phenotypic impact of genomic variation in relevant tissues for human diseases. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=78 SRC="FIGDIR/small/743997v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@ac8319org.highwire.dtl.DTLVardef@1d31920org.highwire.dtl.DTLVardef@169225org.highwire.dtl.DTLVardef@1d4dc02_HPS_FORMAT_FIGEXP M_FIG C_FIG

12
TBpop: an open-access genomic portal integrating genomic variation, population genetic statistics, phylogeny, pangenome composition, and strain metadata of epidemic Mycobacterium tuberculosis strains from China

Zhou, Y.; Huang, F.; Zhao, Y.

2026-08-12 bioinformatics 10.64898/2026.08.06.742672 medRxiv
Top 0.7%
0.8%
Show abstract

Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.

13
Structural and biochemical analysis of the Estrogen-Related Receptor alpha and complex with TMPRSS2 promoter DNA

K, C.; Saxena, A. K.

2026-08-19 cancer biology 10.64898/2026.08.19.744156 medRxiv
Top 0.8%
0.6%
Show abstract

In TMPRSS2 fusion-positive prostate cancer, ERR is involved in regulation of ERG and promotes the androgen receptor independent signaling in the cancer progression. The ERR binds to the ERREs (estrogen-related receptor response elements) present at -5042 bp of the TMPRSS2- promoter and enhances the ERG overexpression that causes prostate cancer progression. To dissect the structural basis of the ERR recognition to the TMPRSS2 promoter DNA, we have purified the full-length ERR (ERRFL), NTD deleted construct (ERR{Delta}NTD), and the DNA-binding domain (ERRDBD) proteins and performed the binding analysis with 30 bp TMPRSS2-promoter DNA (5' -AGTCCAAGGTCGGTGGATC ACAAGGTCAGG-3'). Circular dichroism analysis showed that all three ERR proteins adopt native secondary structures. DNA binding induced subtle changes in the secondary structures, while enhancing the thermal stability (Tm) of all ERRa proteins. Binding analysis showed that ERRDBD bound weakly to the DNA, whereas ERRFL and ERR{Delta}NTD exhibited substantially higher affinities ~120-fold and ~131-fold than ERRaDBD, respectively. Small-angle X-ray scattering (SAXS) analyses revealed a dimeric ERRFL structure and an ERRFL-DNA complex (2:1) structure in solution and fitted well with Alpha Fold model of apo and DNA bound complex of ERRFL. Furthermore, 100 ns dynamics simulations on apo and DNA-bound ERRa proteins showed that all proteins remained structurally stable, with flexibility largely confined to loop regions of ERRa proteins. Our biophysical, DNA binding and structural analyses have revealed the mechanism involved in ERR recognition of the TMPRSS2- promoter DNA, which provides insight into ERR-mediated transcriptional regulation and development of anticancer drugs against ERR-driven prostate cancer.

14
A confound-diagnostic toolkit for in silico perturbation with single-cell foundation models

Qiu, R.; Zhao, M. M.

2026-08-07 bioinformatics 10.64898/2026.08.04.732812 medRxiv
Top 0.8%
0.6%
Show abstract

Deleting a gene token from a cells input sequence offers a convenient native strategy for in silico perturbation, but the resulting embedding delta may not represent a biological knockout response. Apparent effects can instead reflect gene identity, universal responsiveness, limited tokenization coverage, library-size contamination, or circular state scoring. Here, we present a confound-diagnostic framework combining held-out increment testing, responsiveness adjustment, coverage gating, library-size diagnostics, and de-circularized state-shift analysis, together with a numerically matched reimplementation of frozen Geneformers perturbation engine. Across Frangieh and Replogle datasets and linear and nonlinear readouts, the native embedding delta provided no reproducible held-out improvement beyond gene identity. Signal-injection calibration showed that the test detected injected residual signal, whereas native increments remained below its detection floor. Matched controls traced apparent positives to raw-count library-size structure, broad responsiveness, and self-referential scoring, while coverage constrained perturbation applicability and estimate stability without establishing biological specificity. This model-adaptable framework helps determine when foundation-model perturbation readouts warrant biological interpretation. MotivationFoundation-model in silico perturbation could predict perturbation effects when matched experimental data are unavailable. However, in zero-shot settings, embedding-derived responses may reflect gene identity, universal responsiveness, tokenization limits, library-size artifacts, or circular state scoring rather than biological knockout effects. We therefore developed a reusable confound-diagnostic framework that applies matched controls to test whether native perturbation readouts contain information beyond these confounds and warrant biological interpretation.

15
DeMoP: A Language-Model-Guided Mixture-of-Experts Framework for Cancer Prognosis

Tang, C.; Yu, L.; Li, Q.; Xu, L.

2026-08-27 bioinformatics 10.64898/2026.08.24.746579 medRxiv
Top 0.9%
0.6%
Show abstract

Integrating heterogeneous clinical and molecular data for cancer prognosis remains challenging because their dimensionality, semantics and distributions differ across patients and cohorts. Here we present DeMoP, a language-model-guided mixture-of-experts framework that serializes structured patient profiles as natural-language sequences and learns adaptive prognostic representations from clinical variables, copy-number alterations, and gene descriptions. DeMoP combines a fine-tuned DeBERTa-v3-large encoder, attention-based token pooling, and a residual mixture-of-experts prediction head. In held-out tests from two independent pan-cancer cohorts, GENIE (63,090 patients) and TCGA (4,123 patients), DeMoP outperformed the conventional machine-learning and deep-learning baselines evaluated, achieving AUROCs of 0.939 and 0.805 and class-1 F1 scores of 0.72 in both cohorts. A GENIE-trained model transferred directly to TCGA with an overall class-1 F1 score of 0.62. Gene-level ablations recovered established cancer-associated genes and highlighted less-studied candidates. DeMoP provides a unified approach to heterogeneous biomedical data integration, cross-cohort outcome prediction, and model interpretation.

16
Enhancer RNA like function of intergenic inherited lncRNAs during maternal to zygotic transition in zebrafish

Joshi, D. C.; Guha, S.; Ahmed, N.; Dayal, S.; Pillai, B.

2026-08-19 developmental biology 10.64898/2026.08.15.744761 medRxiv
Top 0.9%
0.6%
Show abstract

The maternal-to-zygotic transition (MZT) is a major developmental event during which inherited transcripts are remodeled and zygotic transcription is established. Although parentally inherited long noncoding RNAs (lncRNAs) are present in early embryos, they have been thought to be dispensable. We have identified more than 2000 inherited lncRNAs in zebrafish embryos, but how these RNAs participate in regulatory programs during early development has remained unexplored. Here, the inheritance of selected zebrafish lncRNAs spanning a broad expression range were confirmed at the pre-MZT stage and full-length sequences were captured by Direct RNA nanopore sequencing. We show that 30% inherited intergenic lncRNAs are preferentially associated with active enhancers, annotated as such in DANIO CODE, whereas non-inherited intergenic lncRNAs rarely overlap with enhancers. Perturbation of five inherited intergenic lncRNAs, individually, using antisense oligonucleotides reduced the expression of their respective neighboring genes at 2.5, 4.3, and/or 6 hours post fertilization, indicating that these RNAs act as positive local regulators during MZT. Together, these findings identify inherited intergenic lncRNAs as enhancer-associated regulators with elncRNA-like properties during early embryogenesis.

17
Multiplexed Quantification of Variant Abundance in the Globin Gene Family: Integrating Saturation Mutagenesis with Cross-Paralog Prediction

Cai, X.; Wang, D.; Hu, J.; Huang, Y.; Guo, W.; Shi, Y.; Zhou, Y.; Xiao, C.; Ye, Y.; Wang, C.; Zhou, W.; Xu, X.; Jia, X.

2026-08-24 genetics 10.64898/2026.08.19.745862 medRxiv
Top 1.0%
0.6%
Show abstract

Widespread genetic testing has expanded variant identification, yet functional characterization remains a bottleneck in genome guided medicine. Here, we present a modified Variant Abundance by Massively Parallel Sequencing (VAMP-seq) platform integrating experimental and computational approaches for high-resolution abundance profiling of protein variants. Utilizing a lentiviral integration system, we systematically assessed the stability effects of 2,696 amino acid substitutions in {zeta}-globin (HBZ) via saturation mutagenesis in human cells, achieving complete variant coverage with high reproducibility. Representative variants showed strong concordance with orthogonal low-throughput validation assays. We further developed a deep learning framework leveraging VAMP-seq derived HBZ data to predict variant abundance across thalassemia-associated globin paralogs (HBA, HBB, and HBG1) not experimentally tractable. Our hybrid framework demonstrates how targeted experimental profiling combined with AI-driven extrapolation can accelerate variant interpretation across protein family members.

18
Epromoter 3D interaction-associated regulation in T-acute Lymphoblastic Leukemia

Roca Paixao, J. F.; Manosalva, I.; Pinton, A.; Cieslak, A.; Cardone, C.; Sakakini, N.; Sadouni, N.; Zanzoni, A.; Andrieu, G.; Asnafi, V.; Touzart, A.; Spicuglia, S.

2026-08-26 cancer biology 10.64898/2026.08.25.746923 medRxiv
Top 1%
0.5%
Show abstract

Background: Promoters have been traditionally seen as contiguous gene-adjacent cis-regulatory elements. Yet, substantial studies corroborate that Epromoters (promoters with enhancer activity) engage in distal forms of gene regulation. Although in the three-dimensional (3D) space enhancer-promoter networks have been well studied, the contribution of the circuits of promoter-promoter (P-P) interactions is poorly understood. Furthermore, whether the regulatory aspects of P-P interactions in cancer may be controlled by physical 3D-mediated Epromoter interactions remains elusive. Results: We show that Epromoter-mediated 3D interactions regulate target genes and participate in cluster co-regulation, playing a critical role in T-cell acute Lymphoblastic Leukemia (T-ALL). To achieve this, we first leveraged survival CRISPR screenings in T-ALL model cells (Jurkat) to identify potential Epromoters. By integrating these findings with an H3K27ac HiChIP dataset from T-ALL cells, we characterized a set of Epromoters that establish 3D genome interactions with other promoters. We observed that promoters organize into dense, promoter-rich genomic clusters, and that among them, the clusters enriched with Epromoters actively regulate complex gene expression networks. To investigate gene coregulation, we integrated transcriptomic data from T-ALL patients and found that promoter-promoter (P-P) pairs exhibit positive correlation at multiple levels, and that several Jurkat Epromoter candidate clusters are significantly co-regulated in the patient cohort. To experimentally validate these candidates, we utilized CRISPRi to inhibit Epromoters, which revealed direct transcriptional regulation of multiple target genes within each hub. Finally, we performed cell competition assays to confirm that these Epromoters are vital for T-ALL cell survival. Conclusions: Our analysis provides support for the role of Epromoters in the regulation of 3D P-P interactions and co-regulation of promoter hubs, and how these interactions play a critical part in T-ALL cell survival.

19
Interpretable Forecasting of Kidney Cancer Progression via Generative AI and Symbolic Reasoning

Prol-Castelo, G.; Syrri, E.; Manginas, N.; Manginas, V.; Sanchez-Valle, J.; Katzouris, N.; Paliouras, G.; Valencia, A.; Cirillo, D.

2026-08-26 bioinformatics 10.64898/2026.08.23.746526 medRxiv
Top 1%
0.5%
Show abstract

Predicting cancer stage progression from omics data, and deriving molecular insight into the mechanisms driving it, remains a major challenge, owing in part to the lack of adequate longitudinal data and the interpretability limitations of current forecasting models. Large cancer datasets such as TCGA capture patient profiles cross-sectionally rather than longitudinally, complicating timely treatment decisions as tumors become more invasive. Deep neural networks typically used for forecasting, such as LSTMs, compound this problem by remaining largely opaque and offering clinicians no straightforward way to audit their predictions. Clear cell renal cell carcinoma (ccRCC) illustrates the clinical stakes of both challenges. Five-year survival falls from over 94% at stage I to 28% at stage IV, yet early-stage tumors are often managed under active surveillance, a strategy constrained by sparse molecular evidence of progression risk. Detecting progression in time, meanwhile, demands forecasts clinicians can interpret and trust, not black-box predictions. We address both challenges by combining generative and symbolic AI: a Variational Autoencoder trained on bulk RNA-Seq profiles of 530 TCGA ccRCC patients generates synthetic pseudo-time trajectories that overcome the absence of longitudinal data, while a symbolic rule-induction framework (ASAL) learns finite-state automata from these trajectories, encoding stage transition as human-readable Boolean conditions over gene expression, which a complex event forecasting system (Wayeb) converts into probabilistic forecasts of stage advancement. An independent XGBoost classifier trained on real patients (F1 score = 0.71-0.81) shows a gradual early-to-late probability shift along the synthetic trajectories, absent in non-progressing control trajectories. Pathway enrichment of those trajectories reveals stage-dependent changes in established kidney cancer-related processes, including the TCA cycle and DNA repair. Finally, our symbolic forecaster nearly matches an LSTM baseline (macro F1 = 0.928 vs. 0.964), while additionally offering an inspectable rule set and a probability distribution over transition timing rather than a single opaque score. This work shows that generative and symbolic AI, paired together, can turn cross-sectional cohorts into a transparent, forecast-oriented framework for modeling disease progression, demonstrated here in ccRCC.

20
A Conversational Multi-Agent AI System for Integrated Multi-Omics Analysis and Biomedical Discovery

Rajdeo, P.; Asanuma, S.; Kouril, M.; Lu, P.; Chen, J.; Chadha, A.; Prasath, V. B. S.; Aronow, B. J.; Salomonis, N.

2026-08-14 bioinformatics 10.64898/2026.08.08.743577 medRxiv
Top 1%
0.5%
Show abstract

Single-cell and spatial omics offer unprecedented opportunities to decipher the mechanisms of disease, however, this process requires teams of experts, iterative trial-and-error and reasoning across modalities. Here we present LungChat (https://chat.lungmap.net), a conversational system for integrated multi-omics analysis and biomedical discovery, deployed as a hierarchical multi-agent architecture in which a supervisor decomposes natural-language questions into parallel, tool-grounded tasks spanning single-cell and spatial analyses, literature and clinical-trial synthesis, and drug repurposing. To predict new therapeutics, LungChat implements Direction-Aware Repurposing and Targeting (DART) to distinguish perturbations that reverse disease transcriptional programs from those that reinforce them, at the cell-type level, for safety prediction. Controlled architecture ablations showed that hierarchical orchestration improved grounded abstention and token efficiency and preserved strong performance on complex multi-step tasks. In pulmonary disease case studies, LungChat independently prioritized saracatinib for IPF through drug-connectivity screening, followed by DART-based cell-type analysis; the same compound has been evaluated in the STOP-IPF clinical trial (NCT04598919). The system also recovered fluticasone propionate, an established COPD therapy, through a single orchestrated analysis. This tissue-agnostic system provides a blueprint for verifiable agentic AI systems that support reproducible scientific discovery.